BMC Medical Research Methodology
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match BMC Medical Research Methodology's content profile, based on 47 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Curnow, E.; Carpenter, J. R.; Heron, J. E.; Cornish, R. P.; Rach, S.; Didelez, V.; Langeheine, M.; Tilling, K.
Show abstract
BackgroundEpidemiological studies often have missing data. Multiple imputation (MI) is a commonly-used strategy for such studies. MI guidelines for structuring the imputation model have focused on compatibility with the analysis model, but not on the need for the (compatible) imputation model(s) to be correctly specified. Standard (default) MI procedures use simple linear functions. We examine the bias this causes and performance of methods to identify problematic imputation models, providing practical guidance for researchers. MethodsBy simulation and real data analysis, we investigated how imputation model mis-specification affected MI performance, comparing results with complete records analysis (CRA). We considered scenarios in which imputation model mis-specification occurred because (i) the analysis model was mis-specified, or (ii) the relationship between exposure and confounder was mis-specified. ResultsMis-specification of the relationship between outcome and exposure, or between exposure and confounder in the imputation model for the exposure, could result in substantial bias in CRA and MI estimates (in addition to any bias in the full-data estimate due to analysis model mis-specification). MI by predictive mean matching could mitigate for model mis-specification. Model mis-specification tests were effective in identifying mis-specified relationships. These could be easily applied in any setting in which CRA was, in principle, valid and data were missing at random (MAR). ConclusionWhen using MI methods that assume data are MAR, compatibility between the analysis and imputation models is necessary, but is not sufficient to avoid bias. We propose an easy-to-follow, step-by-step procedure for identifying and correcting mis-specification of imputation models.
Lusa, L.; Kappenberg, F.; Collins, G. S.; Schmid, M.; Sauerbrei, W.; Rahnenfuehrer, J.
Show abstract
The number of prediction models proposed in the biomedical literature has been growing year on year. In the last few years there has been an increasing attention to the changes occurring in the prediction modeling landscape. It is suggested that machine learning techniques are becoming more popular to develop prediction models to exploit complex data structures, higher-dimensional predictor spaces, very large number of participants, heterogeneous subgroups, with the ability to capture higher-order interactions. We examine these changes in modelling practices by investigating a selection of systematic reviews on prediction models published in the biomedical literature. We selected systematic reviews published since 2020 which included at least 50 prediction models. Information was extracted guided by the CHARMS checklist. Time trends were explored using the models published since 2005. We identified 8 reviews, which included 1448 prediction models published in 887 papers. The average number of study participants and outcome events increased considerably between 2015 and 2019, but remained stable afterwards. The number of candidate and final predictors did not noticeably increase over the study period, with a few recent studies using very large numbers of predictors. Internal validation and reporting of discrimination measures became more common, but assessing calibration and carrying out external validation were less common. Information about missing values was not reported in about half of the papers, however the use of imputation methods increased. There was no sign of an increase in using of machine learning methods. Overall, most of the findings were heterogeneous across reviews. Our findings indicate that changes in the prediction modeling landscape in biomedicine are less dramatic than expected and that poor reporting is still common; adherence to well established best practice recommendations from the traditional biostatistics literature is still needed. For machine learning best practice recommendations are still missing, whereas such recommendations are available in the traditional biostatistics literature, but adherence is still inadequate.
Guski, J.; Aborageh, M.; Fröhlich, H.
Show abstract
BackgroundTargeted estimation offers a robust and unbiased approach for causal inference of the average treatment effect (ATE) from observational data, even with confounding, dependent censoring, and competing risks. Its advantages include double robustness, statistical rigor, and flexible data-adaptive modeling, potentially leveraging machine/deep learning. However, existing implementations lack model selection flexibility and are R-based, hindering adoption by the Python-focused machine learning community. ResultsWe propose PyTMLE, a flexible Python package for causal machine learning-based targeted estimation with survival outcomes and competing risks. PyTMLE supports scikit-survival and pycox, and inbuilt robustness checks based on E-values. PyTMLE is easy to use with initial estimation of nuisance parameters that are obtained via super learning by default. We showcase its basic usage on the established Hodgkins disease dataset, where our package reveals the protective effect of chemotherapy on relapse risk. ConclusionsThis package promotes targeted estimation in time-to-event analysis for applied machine learning, enabling fully data-adaptive nuisance parameter estimation, potentially with deep learning. Future enhancements may include time-dependent confounders and dynamic treatment regimes.
DeVito, N. J.; Bacon, S.; Goldacre, B.
Show abstract
IntroductionNon-publication of clinical trials results is an ongoing issue. In 2016 the US government updated the results reporting requirements to ClinicalTrials.gov for trials covered under the FDA Amendments Act 2007. We set out to develop and deliver an online tool which publicly monitors compliance with these reporting requirements, facilitates open public audit, and promotes accountability. MethodsWe conducted a review of the relevant legislation to extract the requirements on reporting results. Specific areas of the statutes were operationalized in code based on the results of our policy review, publicly available data from ClinicalTrials.gov, and communications with ClinicalTrials.gov staff. We developed methods to identify trials required to report results, using publicly available registry data; to incorporate additional relevant information such as key dates and trial sponsors; and to determine when each trial became due. This data was then used to construct a live tracking website. ResultsThere were a number of administrative and technical hurdles to successful operationalization of our tracker. Decisions and assumptions related to overcoming these issues are detailed along with clarifications directly from ClinicalTrials.gov. The FDAAA TrialsTracker was successfully launched in February 2018 and provides users with an overview of results reporting compliance. DiscussionClinical trials continue to go unreported despite numerous guidelines, commitments, and legal frameworks intended to address this issue. In the absence of formal sanctions from the FDA and others, we argue tools such as ours - providing live data on trial reporting - can improve accountability and performance. In addition, our service helps sponsors identify their own individual trials that have not yet reported results: we therefore offer positive practical support for sponsors who wish to ensure that all their completed trials have reported.
Mullaert, J.; Schmeller, S.; Austin, P. C.; Latouche, A.
Show abstract
When fitting competing risks regression models, a variety of variable selection methods exist, including backward selection on the subdistribution hazard, on the cause-specific hazards, and penalized methods. However, a benchmark study comparing these different procedures is lacking. We conducted an extensive simulation study to compare three variable selection procedures in terms of both model selection ability and predictive accuracy. 5120 datasets were simulated in various conditions aiming at being representative of real applications in clinical epidemiology. Results show that the backward selection procedure can lead to high false discovery rate (FDR) because of implementation choices. Even for scenarios with a high numbers of events per variable (EPV), the true model is rarely identified by any of the tested procedures. Survival predictions were assessed with time-dependent AUC and show similar performances for all methods. We also provide an application on real data from stem cell transplanted patients in hematology. We conclude that the identification of the true model in competing risk regression is a very difficult task, and suggest some recommendations to analysts: (1) to report event per variable for the event type of interest and (2) to use multiple methods to deal with model uncertainty and avoid implementation pitfalls.
Köster, D.; Chaturvedi, M.; Rübsamen, N.; Bapda, M.; Karch, A.; Zapf, A.
Show abstract
BackgroundDuring epidemics with emerging infections, diagnostic tests directly inform model-based decision-making and thereby shape infection control strategies. However, diagnostic accuracy studies (DTA) assessing the validity of these tests must be conducted under severe time and data constraints. We investigated whether the integration of adaptive designs and of epidemic spread modelling for prevalence prediction can accelerate DTA studies during epidemics with emerging infections without compromising statistical validity. MethodsWe compared three designs in a large-scale simulation study using a COVID-19 use case: a fixed design; a standard adaptive design with unblinded interim analysis enabling early stopping or sample size adaptation; and an adaptive design additionally integrating a prevalence projection model to inform sample size re-estimation. Data-generating mechanisms were based on infectious disease models and realistic recruitment constraints. As decision rules we used in one simulation line WHO criteria for DTA studies for COVID-19 and in the other one more liberal performance thresholds. Across 1,440 factorial scenarios (5,000 replications each), we evaluated study duration, sample size requirements, statistical power as well as bias in estimates. ResultsBoth adaptive designs enabled substantial operational gains. For the WHO thresholds, early stopping (for futility or infeasibility) occurred in 80% of adaptive simulations; early efficacy stops were rare. Under more liberal thresholds, early termination was less frequent, leading to more studies reaching final analysis. Required sample sizes under WHO criteria frequently exceeded 10,000 participants, making fixed designs practically infeasible. Adaptive designs identified infeasible scenarios early and avoided continuation. Under liberal thresholds, recalculated sample sizes in adaptive designs closely tracked theoretical needs up to the upper quartile, in contrast to fixed designs mirroring the low power commonly observed in real-world pandemic studies. Overall, adaptive designs shortened study duration when stopping early and prevented continuation of unpromising trials. DiscussionAdaptive designs in DTA studies during epidemics with emerging pathogens improve feasibility by preventing unrealistic recruitment targets and enabling early abandonment of non-viable scenarios. When realistic performance thresholds are used, adaptive re-estimation produces sample sizes more aligned with statistical requirements without systematic operational penalties. These findings support the adoption of adaptive approaches in confirmatory DTA studies for emerging infections as a pragmatic response to time pressure and uncertainty.
Lusa, L.; Proust-Lima, C.; Schmidt, C. O.; Lee, K. J.; Baillie, M.; le Cessie, S.; Frank, L.; Huebner, M.
Show abstract
Initial data analysis (IDA) is the part of the data pipeline that takes place between the end of data retrieval and the beginning of data analysis that addresses the research question. Systematic IDA and clear reporting of the IDA findings is an important step towards reproducible research. A general framework of IDA for observational studies includes data cleaning, data screening, and possible updates of pre-planned statistical analyses. Longitudinal studies, where participants are observed repeatedly over time, pose additional challenges, as they have special features that should be taken into account in the IDA steps before addressing the research question. We propose a systematic approach in longitudinal studies to examine data properties prior to conducting planned statistical analyses. In this paper we focus on the data screening element of IDA, assuming that the research aims are accompanied by an analysis plan, meta-data are well documented, and data cleaning has already been performed. IDA screening domains are participation profiles over time, missing data, and univariate and multivariate descriptions, and longitudinal aspects. Executing the IDA plan will result in an IDA report to inform data analysts about data properties and possible implications for the analysis plan that are other elements of the IDA framework. Our framework is illustrated focusing on hand grip strength outcome data from a data collection across several waves in a complex survey. We provide reproducible R code on a public repository, presenting a detailed data screening plan for the investigation of the average rate of age-associated decline of grip strength. With our checklist and reproducible R code we provide data analysts a framework to work with longitudinal data in an informed way, enhancing the reproducibility and validity of their work.
Ouedraogo, F. A. S.
Show abstract
Despite the evolution of epidemiological analysis and modeling tools, difficulties still remain, especially in developing countries, regarding the availability and use of these tools. Often expensive, requiring high technical expertise, demanding constant connectivity of several or sometimes even significant resources, these tools, although efficient, present a major gap with the operational realities of health districts. It is in this context that we introduce Episia, an open-source Python library designed and conceived to provide a framework to facilitate epidemiological analysis and modeling. It integrates a suite of compartmental epidemic models (SIR, SEIR, SEIRD) with a sensitivity analysis using the Monte Carlo method, a complete biostatistics suite validated against the OpenEpi reference standard, as well as a native DHIS2 client for automated data ingestion. Developed in Burkina Faso, it is optimized and aims not only to address these health challenges encountered in Africa but also remains a versatile tool for global health informatics.
GUO, Y.; STRAUSS, V. Y.; PRIETO-ALHAMBRA, D.; Khalid, S.
Show abstract
BackgroundThe surge of treatments for COVID-19 in the ongoing pandemic presents an exemplar scenario with low prevalence of a given treatment and high outcome risk. Motivated by that, we conducted a simulation study for treatment effect estimation in such scenarios. We compared the performance of two methods for addressing confounding during the process of estimating treatment effects, namely disease risk scores (DRS) and propensity scores (PS) using different machine learning algorithms. MethodsMonte Carlo simulated data with 25 different scenarios of treatment prevalence, outcome risk, data complexity, and sample size were created. PS and DRS matching with 1: 1 ratio were applied with logistic regression with least absolute shrinkage and selection operator (LASSO) regularization, multilayer perceptron (MLP), and eXtreme Gradient Boosting (XgBoost). Estimation performance was evaluated using relative bias and corresponding confidence intervals. ResultsBias in treatment effect estimation increased with decreasing treatment prevalence regardless of matching method. DRS resulted in lower bias compared to PS when treatment prevalence was less than 10%, under strong confounding and nonlinear nonadditive data setting. However, DRS did not outperform PS under linear data setting and small sample size, even when the treatment prevalence was less than 10%. PS had a comparable or lower bias to DRS when treatment prevalence was common or high (10% - 50%). All three machine learning methods had similar performance, with LASSO and XgBoost yielding the lowest bias in some scenarios. Decreasing sample size or adding nonlinearity and non-additivity in data improved the performance of both PS and DRS. ConclusionsUnder strong confounding with large sample size DRS reduced bias compared to PS in scenarios with low treatment prevalence (less than 10%), whilst PS was preferable for the study of treatments with prevalence greater than 10%, regardless of the outcome prevalence. Key MessagesO_LIWhen handling nonlinear nonadditive data with strong confounding, DRS estimated by machine learning methods outperforms PS in scenarios with low treatment prevalence (less than 10%). C_LIO_LIHowever, if having linear data and small sample size data with strong confounding, we did not observe DRS outperformed PS even when treatment prevalence was less than 10%. C_LIO_LIOur results suggested that using PS performed better compared to DRS in tackling strong confounding problems with treatment prevalence greater than 10%. C_LIO_LISmall sample size increased bias for both DRS and PS methods, and it affected DRS more than PS. C_LI
ORWA, F. O.; Mutai, C.; Nizeyimana, I.; Mwangi, A.
Show abstract
When randomized controlled trials are impractical, interrupted time series designs offer a rigorous quasi-experimental approach to assess population level policies. Indeed, in the context of quasi-experimental designs (QEDs), the Interrupted Time Series (ITS) method is commonly thought of as the most robust. But interrupted time series designs are susceptible to serial correlation and confounding by time-varying factors associated with both the intervention and the outcome, which may result in biased inference. Thus, we provide a simulation-based contrast of controlled interrupted time series (CITS) and multivariable regression (multivariable negative binomial regression) for estimation of policy effects in count time series data. These approaches are widely used in policy evaluations, yet their comparative performance in typical population health settings has rarely been examined directly. We tested both approaches within a variety of data generating situations, differing in the series length, intervention effect size, and magnitude of lag-1 autocorrelation. Bias, standard error calibration, confidence interval coverage, mean squared error, and statistical power were assessed for performance. Both methods gave unbiased estimates for moderate and large intervention effects, although bias was more pronounced for small effects, particularly in short series. Although the point estimate performance was similar, inferential properties varied significantly. CITS always had smaller mean squared error, better consistency between model based and empirical standard errors, and confidence interval coverage near the 95% nominal levels over weak to moderate autocorrelation. By contrast, multivariable regression was more sensitive to serial dependence, leading to underestimated standard errors and undercoverage, especially at moderate to high autocorrelation, regardless of Newey-West adjustments. These findings show the benefits of using a concurrent control series and the importance of structurally accounting for serial correlation when studying population level policies with time series data.
Blagus, R.; Leskosek, B.; Ortega, F. B.; Tomkinson, G. R.; Jurak, G.
Show abstract
Norm-referenced tests compare individuals to a group. While norms are often presented in tables and graphs, exact score evaluation relies on model parameters, often undisclosed. These models, like those from the R gamlss package, include individual data protected by law and consent, hindering full transparency. Thus, this paper proposes standards for publishing test norms that allow precise score evaluation while protecting participant privacy. We outline specific requirements for norms publications: a) the exact presentation of the fitted model that contains the estimates of all model parameters and other information required for exact evaluation; b) computer sharable fit of the model that does not contain any sensitive information and can be used by those with programming skills to evaluate scores; and c) a web-based application that can be used by those without programming skills to use the results of the fitted model. To facilitate publication and utilization of norms, we have developed and provided in this manuscript an open-source R package of tools for authors and users alike.
Oulhaj, A.; Ahmed, L. A.; Prattes, J.; Suliman, A.; Al Suwaidi, A.; Al-Rifai, R. H.; Sourij, H.; Van Keilegom, I.
Show abstract
BackgroundA plethora of studies on COVID-19 investigating mortality and recovery have used the Cox Proportional Hazards (Cox PH) model without taking into account the presence of competing risks. We investigate, through extensive simulations, the bias in estimating the hazard ratio (HR) and the absolute risk reduction (ARR) of death when competing risks are ignored, and suggest an alternative method. MethodsWe simulated a fictive clinical trial on COVID-19 mimicking studies investigating interventions such as Hydroxychloroquine, Remdesivir, or convalescent plasma. The outcome is time from randomization until death. Six scenarios for the effect of treatment on death and recovery were considered. The HR and the 28-day ARR of death were estimated using the Cox PH and the Fine and Gray (FG) models. Estimates were then compared with the true values, and the magnitude of misestimation was quantified. ResultsThe Cox PH model misestimated the true HR and the 28-day ARR of death in the majority of scenarios. The magnitude of misestimation increased when recovery was faster and/or chance of recovery was higher. In some scenarios, this model has shown harmful treatment effect when it was beneficial. Estimates obtained from FG model were all consistent and showed no misestimation or changes in direction. ConclusionThere is a substantial risk of misleading results in COVID-19 research if recovery and death due to COVID-19 are not considered as competing risk events. We strongly recommend the use of a competing risk approach to re-analyze relevant published data that have used the Cox PH model.
Nyberg, J.; Jonsson, E. N.; Karlsson, M. O.; Häggström, J.
Show abstract
Two full model approaches was compared with respect to their ability to handle missing covariate information. The reference data analysis approach was the full model method in which the covariate effects are estimated conventionally using fixed effects, and missing covariate data is imputed with the median of the non-missing covariate information. This approach was compared to a novel full model method which treats the covariate data as observed data and estimates the covariates as random effects. A consequence of this way of handling the covariates is that no covariate imputation is required and that any missingness in the covariates is handled implicitly. The comparison between the two analysis methods was based on simulated data from a model of height for age z-scores as a function of age. Data was simulated with increasing degrees of randomly missing covariate information (0-90%) and analyzed using each of the two analysis approaches. Not surprisingly, the precision in the parameter estimates from both methods decreased with increasing degrees of missing covariate information. However, while the bias in the parameter estimates increased in a similar fashion for the reference method, the full random effects approach provided unbiased estimates for all degrees of covariate missingness.
Dankwa, E. A.; Cavalli, L.; Balasubramanian, R.; Can, M. H.; Cui, H.; Jia, K. M.; Li, Y.; Ofori, S. K.; Swartwood, N. A.; Wade, C.; Buckee, C. O.; Imai-Eaton, J. W.; Menzies, N. A.
Show abstract
Objective/BackgroundTransmission-dynamic models are commonly used to study infectious disease epidemiology. Calibration involves identifying model parameter values that align model outputs with observed data or other evidence. Inaccurate calibration and inconsistent reporting produce inference errors and limit reproducibility, compromising confidence in modeled results. No standardized framework exists for reporting on calibration of infectious disease models, and an understanding of current calibration approaches is lacking. MethodsWe developed a 15-item framework for reporting calibration practices and applied it in a scoping review to assess calibration approaches and evaluate reporting comprehensiveness in transmission-dynamic models of tuberculosis, HIV and malaria published between January 1, 2018, and January 16, 2024. We searched relevant databases and websites to identify eligible publications, including peer-reviewed studies where these models were calibrated to empirical data or published estimates. ResultsWe identified 411 eligible studies encompassing 419 models, with 74% (n=309) being compartmental models and 20% (n=82) individual-based models (IBMs). The predominant analytical purpose was to evaluate interventions (71% of models, n=298). Parameters were calibrated mainly because they were unknown or ambiguous (40%, n=168), or because determining their value was relevant to the scientific question beyond being necessary to run the model (20%, n=85). The choice of calibration method was significantly associated with model structure (p-value<0.001) and stochasticity (p-value=0.006), with approximate Bayesian computation more frequently used with IBMs and Markov-Chain Monte Carlo with compartmental models. Regarding reporting comprehensiveness, all 15 framework items in the framework were reported in 4% (n=18) of models; 11-14 items in 66% (n=277), and 10 or fewer items in 28% (n= 124). Implementation code was the least reported, available in only 20% (n=82) of models. ConclusionsReporting on calibration is heterogeneous in recent infectious disease modeling literature. Our proposed framework for reporting of calibration approaches could support improved reproducibility and credibility of modeled analyses. Author SummaryCalibration, the identification of parameter values so that model outcomes are consistent with observed data or other evidence, is often employed in the process of obtaining model results to inform health decision making. Despite its importance, there has not been a standardized framework for reporting how calibration is conducted in infectious disease modeling studies. This has led to inconsistent reporting practices and challenges in reproducing model results, potentially compromising confidence in their validity. We developed a calibration reporting framework, based on best practices found in the literature and informed by our expertise in conducting calibration. To assess calibration practices and their reporting, we applied our framework in a scoping review of 419 infectious disease transmission models of HIV, TB and malaria published between 2018 and 2024. Most models reviewed were compartmental (74%) or individual-based (20%), and the choice of calibration methods was associated with model structure and stochasticity. Calibration was conducted predominantly in the context of models aimed at evaluating the impact of disease control interventions, highlighting the role of calibration in decision making. Parameters were calibrated mainly because they were unknown or ambiguous, or because reporting their value was relevant to the scientific question beyond just being necessary to run the model. The comprehensiveness of calibration reporting varied across models, with most models omitting 1 to 5 items in the framework. Accessible implementation code was the most underreported, with only 20% of models including it. Our proposed framework could serve as a tool to standardize calibration reporting, thereby enhancing the transparency and reproducibility of calibration processes in transmission-dynamic models.
Handels, R.; Jonsson, L.; Raket, L. L.; Alzheimer's Disease Neuroimaging Initiative,
Show abstract
INTRODUCTIONRepresentative data of recent Alzheimers Disease (AD) trials are difficult to obtain. We aimed to generate a synthetic version of an original real-world observational dataset, subsequently apply a plausible AD treatment effect, and make our method open-source available. METHODSSynthetic data was generated in the following steps: (1) Obtain real-world data from the ADNI study on demographic (age, sex, education), clinical (cognition: MMSE and ADAS; function: FAQ; composite cognition/function: CDR, ADCOMS) and biological (genetics: APOE4; cerebrospinal fluid: ABeta, Tau; imaging: PET-SUVR-centiloid) outcomes at baseline, 6, 12 and/or 18-month follow-up (35 variables), with missing data multiple-imputed to obtain 10 sets of 537 individuals. (2) Estimate (theoretical) minimum and maximum (all continuous variables) and proportions (all categorical variables). (3) Rescale to 0-1 range (continuous). (4) Estimate beta distribution shape parameters (method of moments; continuous). (5) Transform to cumulative probability distribution function (using shape parameters; continuous) and to cumulative probability (categorical). (6) Transform to a normal distribution. (7) Estimate variance-covariance matrix. (8) Generate random correlated normal data using Cholesky decomposition of variance-covariance. (9) Transform to cumulative probability distribution function. (10) Transform to beta distribution (using shape parameters; continuous). (11) Rescale to original range. (12) Keep half as control arm, and half as intervention arm, and estimate change from baseline. (13) Multiply intervention change from baseline with self-defined hypothetical relative treatment effect. We assumed correlations on normalized scale were similar to correlations on original scale. R code is available on github: https://github.com/ronhandels/synthetic-correlated-data. RESULTSThe synthetic distribution and mean over time showed large similarity to the original data (visually assessed). The absolute difference in pairwise correlations between original and synthetic data median was 0.02 (95th percentile=0.11, max=0.18). CONCLUSIONWe judged our method sufficiently valid to generate synthetic correlated plausible hypothetical trial results.
Muddiman, R.; Aiello Battan, F. I.; Tazare, J.; Schultze, A.; Boland, F.; Perez, T.; Wei, L.; Walsh, M. E.; Moriarty, F.
Show abstract
PurposeSimulation studies are used in pharmacoepidemiology for evaluating inferential methods in a controlled setting, whereby a known data-generating mechanism allows evaluation of the performance of different approaches and assumptions. This study aimed to review simulation studies performed in pharmacoepidemiology. MethodsWe conducted a review of all papers published in the journal of Pharmacoepidemiology and Drug Safety (PDS) over the period 2017 to 2024. We extracted data on study characteristics and key simulation choices such as the type of data generating mechanism used, inferential methods tested and simulation size. ResultsAmong 42 simulation studies included, 34 (81%) were informing comparative effectiveness/safety studies. 22 studies (52%) used simulation in the context of a clinical condition, and 36 (86%) used Monte-Carlo simulation. Inputs not derived from empirical data alone (n=22, 52%) or in combination with real-world data sources (n=19, 45%) were most often used for data generation. The complexity of simulations was often relatively low: although 31 studies (74%) generated data based on other covariates, time-dependent covariates (n=3) and effects (n=4) were rarely implemented. Bias was the most often used performance measure (n=26, 62%), although notably 18 studies (43%) did not report uncertainty in the method. ConclusionSimulations contributed a relatively small number of articles (3.2 % of 1320) to PDS over 2017 to 2024. Greater focus on evaluating methods and inferential approaches, using simulation studies that are appropriately complex given clinical realities may be beneficial to the pharmacoepidemiology field.
Kawabata, E.; Shapland, C. Y.; Palmer, T. M.; Carslake, D.; Tilling, K.; Hughes, R.
Show abstract
BackgroundUnmeasured confounding is a persistent concern in observational studies. We can quantitatively assess the impact of unmeasured confounding using a quantitative bias analysis (QBA). A QBA specifies the relationship between the unmeasured confounder(s), U, and study data via its bias parameters. There are two broad classes of QBA methods: deterministic and probabilistic. We focus on a probabilistic QBA which incorporates external information about U via prior distribution(s) placed on these bias parameters and can be implemented as a Bayesian QBA or a Monte Carlo QBA. A Bayesian QBA combines the prior distribution(s) with the datas likelihood function whilst a Monte Carlo QBA samples the bias parameters directly from their prior distributions. Software implementations of probabilistic QBAs to unmeasured confounding are scarce and mainly limited to unadjusted analyses of a binary exposure and outcome. One exception is R package unmconf (Hebdon et al 2024, BMC Med. Res. Methodol., https://doi.org/10.1186/s12874-024-02322-2) which implements a Bayesian QBA, applicable when the analysis is a generalised linear model (GLM). However, for a study with q measured confounders and a single U, unmconf requires information on at least 3+ q bias parameters, which is burdensome when q >1 and validation data are unavailable. AimWe propose a flexible Monte Carlo QBA where the number of bias parameters is independent of the number of measured confounders. It is applicable to a GLM or survival proportional hazards model, with binary, continuous, or categorical exposure and measured confounders, and one or multiple ([≥] 2) binary or continuous unmeasured confounders. MethodsVia simulations, we evaluated our Monte Carlo QBA for different analyses (e.g., varying the regression model, type of variables for the exposure and unmeasured confounder), and different levels of dependency between the measured and unmeasured confounders. Also, using our proposed bias model, we compare a Monte Carlo implementation to a fully Bayesian implementation when the analysis is a linear or logistic regression. We repeat the simulation study for prior distributions with different levels of informativeness. ResultsIgnoring U resulted in substantially biased estimates with substantial confidence interval undercoverage (e.g., 57%). Our Monte Carlo QBA (with informative priors) resulted in unbiased (or minimally biased) point estimates and interval estimates with close to nominal coverage. For binary U, levels of bias were marginally higher when U was strongly correlated with the measured confounders. The performances of the Monte Carlo and Bayesian implementations were comparable. ConclusionWe have proposed a flexible probabilistic QBA for unmeasured confounding which is applicable for a wide range of regression-based analyses. We have minimised the burden placed on the user by limiting the number of bias parameters and avoiding the need for specialist knowledge about Bayesian inference or Bayesian software. Our proposed Monte Carlo QBA will be implemented as Stata command and R package, qbaconfound.
Thiesmeier, R.; Madley-Dowd, P.; Ahlqvist, V.; Orsini, N.
Show abstract
IntroductionSystematically missing covariates are a common challenge in medical research synthesis of quantitative data, particularly when individual participant data cannot be shared across study sites. Imputing covariate values in studies where they are systematically unobserved using information from sites where the covariate is observed implicitly assumes similarity of associations across studies. The behaviour of this assumption, and the bias arising from violating it, remains difficult to qualitatively reason about. Here, we evaluated a two-stage imputation approach for handling systematically missing covariates using simulations across a range of statistical and causal heterogeneity scenarios. MethodsWe conducted a simulation study with varying degrees of between-study heterogeneity and systematic differences in model parameters. A binary confounder was set to systematically missing in half of the studies. Study-specific effect estimates were combined using a two-stage meta-analytic model. The performance of the imputation approach was evaluated with the primary estimand being the pooled conditional confounding-adjusted exposure effect across all studies. ResultsBias in the pooled adjusted effect estimate was small across scenarios with low to substantial between-study heterogeneity. Bias increased monotonically with increasingly pronounced differences in causal structures across study sites. Coverage remained close to the nominal level under low to substantial between-study heterogeneity, but deteriorated markedly as differences in causal structures between study sites became more severe. ConclusionThe two-stage cross-site imputation approach produced valid pooled effect estimates across a wide range of simulated scenarios but showed monotonic sensitivity to differences in causal structures across studies. The results provide insight into the conditions under which cross-site imputation may be appropriate for handling systematically missing covariates in research synthesis.
Gillette, J. J.; Lu, M.; Tsybulnik, D. Y.; Heston, T. F.
Show abstract
The fragility quotient normalizes the fragility index by total sample size but can obscure arm-specific fragility when trial arms are unequal, since the fragility index systematically alters outcomes in only one arm. We defined the modified-arm fragility quotient as the fragility index divided by the size of the arm in which outcomes are toggled. We compared it to the fragility quotient using three approaches: (1) re-analysis of 29 categorical outcomes from 18 large-vessel vasculitis trials (allocation ratios 1.0-2.0), (2) R grid simulation across 1:1 to 4:1 allocation ratios with sample sizes 60-240 (90,000 simulations), and (3) Python Monte Carlo simulation with 3,000 replications across 1:1, 1:2, and 1:3 allocations. In balanced 1:1 trials, the modified-arm fragility quotient equaled twice the fragility quotient exactly, as predicted. The proportion of trials in which the modified-arm fragility quotient exceeded twice the fragility quotient increased systematically with allocation imbalance: 0% (1:1), 81% (2:1), 93% (3:1), and 95% (4:1). In the R simulation, mean divergence between the modified-arm fragility quotient and twice the value of the fragility quotient increased from 0.00 (1:1) to 0.10 (4:1). Empirical data showed a strong correlation (r = 0.98) between the two quotients, but divergence increased in the 25 unbalanced allocation trials. The ratio of the modified-arm fragility quotient to twice the value of the fragility quotient scaled linearly with allocation ratio, with mean values of 1.0 (1:1), 1.36 (2:1), 1.91 (3:1), and 2.41 (4:1). These findings show that the modified-arm fragility quotient and total-sample fragility quotient convey equivalent information under equal allocation. In contrast, under unequal randomization, the total-sample fragility quotient can mischaracterize trial fragility due to dilution from the larger, unmodified arm, whereas the modified-arm fragility quotient avoids this bias by restricting normalization to the modified arm.
Abdelhay, O.; Shatnawi, A.; Najadat, H.; Altamimi, T.
Show abstract
IntroductionClass imbalance--situations where clinically important "positive" cases form <30 % of the dataset--systematically degrades the sensitivity and fairness of medical prediction models. Although data-level techniques such as random oversampling, random undersampling and SMOTE, and algorithm-level approaches like cost-sensitive learning, are widely used, the empirical evidence describing when these corrections improve model performance remains fragmented across diseases and modelling frameworks. This protocol outlines a scoping systematic review with meta-regression that will map and quantitatively summarise 15 years of research on resampling strategies in imbalanced clinical datasets, addressing a critical methodological gap in trustworthy medical AI. Methods and analysisWe will search MEDLINE, EMBASE, Scopus, Web of Science Core Collection and IEEE Xplore, plus grey-literature sources (medRxiv, arXiv, bioRxiv) for primary studies (2009 - 31 Dec 2024) that apply at least one resampling or cost-sensitive method to binary clinical prediction tasks with a minority-class prevalence <30 %. No language restrictions will be applied. Two reviewers will screen records, extract data with a piloted form and document the process in a PRISMA flow diagram. A descriptive synthesis will catalogue clinical domain, sample size, imbalance ratio, resampling technique, model type and performance metrics where[≥]10 studies report compatible AUCs, a random-effects mixed-effects meta-regression (logit-transformed AUC) will examine moderators including imbalance ratio, resampling class, model family and sample size. Small-study effects will be probed with funnel plots, Eggers test, trim-and-fill and weight-function models; influence diagnostics and leave-one-out analyses will assess robustness. Because this is a methodological review, formal clinical risk-of-bias tools are optional; instead, design-level screening, influence diagnostics and sensitivity analyses will ensure transparency. DiscussionBy combining a broad conceptual map with quantitative estimates, this review will establish when data-level versus algorithm-level balancing yields genuine improvements in discrimination, calibration and cost-sensitive metrics across diverse medical domains. The findings will guide researchers in choosing parsimonious, evidence-based imbalance corrections, inform journal and regulatory reporting standards, and highlight research gaps, such as the under-reporting of calibration and misclassification costs, that must be addressed before balanced models can be trusted in clinical practice. Systematic review registrationINPLASY202550026